agora inbox for pgsql-hackers@postgresql.orghelp / color / mirror / Atom feed
[PATCH] Lock upgrade without deadlocks. 956+ messages / 2 participants [nested] [flat]
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH] Lock upgrade without deadlocks. @ 2026-04-17 13:13 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 956+ messages in thread From: Antonin Houska @ 2026-04-17 13:13 UTC (permalink / raw) For REPACK (CONCURRENTLY), it's essential to upgrade ShareUpdateExclusive lock to AccessExclusiveLock for the final stage of table processing. However, we cannot release the earlier before requesting the latter, because if another process ran DDL on the table in between, REPACK would have to cancel its transaction, and waste possibly a lot of work. Without releasing ShareUpdateExclusive lock, the upgrade can cause deadlock, and it's not really deterministic whether REPACK or the other process will get canceled. This patch adds a new field to the LOCK structure which tells which process is trying to upgrade the lock - this field is set right before the upgrade. All processes waiting in the lock queue are then woken up to check the flag. When another process tries to get the lock after that, it checks this field, and if it's set, it checks for deadlocks even if deadlock timeout hasn't expired yet. If the process is already in the lock's queue and sleeping, the lock upgrading process wakes it up so it checks the flag immediately. At the time the upgrading process starts to sleep, all the other process should be aware that they need to check for deadlock during their lock acquisitions, so the upgrading process does not have to check for deadlocks when it gets woken up. Thus it should never receive deadlock error report. --- src/backend/commands/repack.c | 4 +- src/backend/storage/lmgr/lmgr.c | 35 ++++++++--- src/backend/storage/lmgr/lock.c | 107 +++++++++++++++++++++++++++++++- src/backend/storage/lmgr/proc.c | 32 +++++++--- src/include/storage/lmgr.h | 1 + src/include/storage/lock.h | 5 +- src/include/storage/proc.h | 4 +- 7 files changed, 166 insertions(+), 22 deletions(-) diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 58e3867246f..310b2a65099 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -3055,8 +3055,10 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, /* * Acquire AccessExclusiveLock on the table, its TOAST relation (if there * is one), all its indexes, so that we can swap the files. + * + * TODO The same for indexes and TOAST? */ - LockRelationOid(old_table_oid, AccessExclusiveLock); + LockRelationOidUpgrade(old_table_oid, AccessExclusiveLock); /* * Lock all indexes now, not only the clustering one: all indexes need to diff --git a/src/backend/storage/lmgr/lmgr.c b/src/backend/storage/lmgr/lmgr.c index 2ccf7237fee..0b103485d54 100644 --- a/src/backend/storage/lmgr/lmgr.c +++ b/src/backend/storage/lmgr/lmgr.c @@ -58,6 +58,8 @@ typedef struct XactLockTableWaitInfo const ItemPointerData *ctid; } XactLockTableWaitInfo; +static void LockRelationOidCommon(Oid relid, LOCKMODE lockmode, + bool isUpgrade); static void XactLockTableWaitErrorCb(void *arg); /* @@ -105,6 +107,21 @@ SetLocktagRelationOid(LOCKTAG *tag, Oid relid) */ void LockRelationOid(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, false); +} + +/* + * Like above, but upgrade an existing lock. + */ +void +LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode) +{ + LockRelationOidCommon(relid, lockmode, true); +} + +static void +LockRelationOidCommon(Oid relid, LOCKMODE lockmode, bool isUpgrade) { LOCKTAG tag; LOCALLOCK *locallock; @@ -113,7 +130,7 @@ LockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, isUpgrade); /* * Now that we have the lock, check for invalidation messages, so that we @@ -157,7 +174,7 @@ ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode) SetLocktagRelationOid(&tag, relid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -191,7 +208,7 @@ LockRelationId(LockRelId *relid, LOCKMODE lockmode) SET_LOCKTAG_RELATION(tag, relid->dbId, relid->relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -254,7 +271,7 @@ LockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, false, true, &locallock, - false); + false, false); /* * Now that we have the lock, check for invalidation messages; see notes @@ -286,7 +303,7 @@ ConditionalLockRelation(Relation relation, LOCKMODE lockmode) relation->rd_lockInfo.lockRelId.relId); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -591,7 +608,7 @@ ConditionalLockTuple(Relation relation, const ItemPointerData *tid, LOCKMODE loc ItemPointerGetOffsetNumber(tid)); return (LockAcquireExtended(&tag, lockmode, false, true, true, NULL, - logLockFailure) != LOCKACQUIRE_NOT_AVAIL); + logLockFailure, false) != LOCKACQUIRE_NOT_AVAIL); } /* @@ -749,7 +766,7 @@ ConditionalXactLockTableWait(TransactionId xid, bool logLockFailure) SET_LOCKTAG_TRANSACTION(tag, xid); if (LockAcquireExtended(&tag, ShareLock, false, true, true, NULL, - logLockFailure) + logLockFailure, false) == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1043,7 +1060,7 @@ ConditionalLockDatabaseObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; @@ -1123,7 +1140,7 @@ ConditionalLockSharedObject(Oid classid, Oid objid, uint16 objsubid, objsubid); res = LockAcquireExtended(&tag, lockmode, false, true, true, &locallock, - false); + false, false); if (res == LOCKACQUIRE_NOT_AVAIL) return false; diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c index c221fe96889..bcbbf36c091 100644 --- a/src/backend/storage/lmgr/lock.c +++ b/src/backend/storage/lmgr/lock.c @@ -810,7 +810,7 @@ LockAcquire(const LOCKTAG *locktag, bool dontWait) { return LockAcquireExtended(locktag, lockmode, sessionLock, dontWait, - true, NULL, false); + true, NULL, false, false); } /* @@ -829,6 +829,9 @@ LockAcquire(const LOCKTAG *locktag, * * logLockFailure indicates whether to log details when a lock acquisition * fails with dontWait = true. + * + * isUpgrade should be true if the backend already holds this lock in a lower + * mode. */ LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, @@ -837,7 +840,8 @@ LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure) + bool logLockFailure, + bool isUpgrade) { LOCKMETHODID lockmethodid = locktag->locktag_lockmethodid; LockMethod lockMethodTable; @@ -1119,7 +1123,8 @@ LockAcquireExtended(const LOCKTAG *locktag, * case, because JoinWaitQueue() may discover that we can acquire the * lock immediately after all. */ - waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait); + waitResult = JoinWaitQueue(locallock, lockMethodTable, dontWait, + isUpgrade); } if (waitResult == PROC_WAIT_STATUS_ERROR) @@ -1225,8 +1230,85 @@ LockAcquireExtended(const LOCKTAG *locktag, Assert(!dontWait); PROCLOCK_PRINT("LockAcquire: sleeping on lock", proclock); LOCK_PRINT("LockAcquire: sleeping on lock", lock, lockmode); + + /* + * Lock upgrade can introduce deadlock, therefore enforce special + * behavior of other processes that deal with this lock. In + * particular, any waiter that sees upgradedBy set is expected to + * perform deadlock check as soon as it's woken up. (Since we are + * already in the queue, the deadlock detector has all the information + * it needs.) + */ + if (isUpgrade) + { + dlist_iter iter; + + /* + * There should not be multiple upgrades at the same time. XXX + * ERROR ? + */ + Assert(lock->upgradedBy == NULL); + + lock->upgradedBy = MyProc; + + /* + * Now, before I sleep myself, wake up all the existing waiters + * (except for me), so they check for deadlock. + */ + dclist_foreach(iter, &lock->waitProcs) + { + PGPROC *proc = dlist_container(PGPROC, waitLink, + iter.cur); + + if (proc != MyProc) + SetLatch(&proc->procLatch); + } + } LWLockRelease(partitionLock); + if (!isUpgrade) + { + HASH_SEQ_STATUS status; + LOCALLOCK *llock; + bool check_deadlock = false; + + /* + * Likewise, new waiters should check the isUpgrade flag and + * engage deadlock detector if needed. The lock being acquired + * should already be in the hash. + * + * Note that, to make deadlock detection happen as soon as + * possible (i.e. regardless deadlock timeout), we check the other + * locks too. XXX There might be ways to get into the deadlock + * indirectly, but that needs more analysis. As long as the + * upgrading process skips deadlock detection altogether, the + * worst consequence of missing some lock wait cycle here is that + * the other process will be kicked-off no sooner than after its + * deadlock timeout has elapsed. (Which in turn means that the + * lock upgrade will take longer than expected.) + */ + hash_seq_init(&status, LockMethodLocalHash); + while ((llock = (LOCALLOCK *) hash_seq_search(&status)) != NULL) + { + LOCK *mylock = llock->lock; + + if (mylock && mylock->upgradedBy) + { + check_deadlock = true; + break; + } + } + /* Check for deadlock if needed. */ + if (check_deadlock) + { + DeadLockState deadlock_state; + + deadlock_state = CheckDeadLock(); + if (deadlock_state == DS_HARD_DEADLOCK) + DeadLockReport(); + } + } + waitResult = WaitOnLock(locallock, owner); /* @@ -1245,6 +1327,18 @@ LockAcquireExtended(const LOCKTAG *locktag, DeadLockReport(); /* DeadLockReport() will not return */ } + + /* + * If finishing the lock upgrade, we should not get into a deadlock + * anymore, so let others know that they do not have to care either. + */ + if (isUpgrade) + { + LWLockAcquire(partitionLock, LW_EXCLUSIVE); + Assert(lock->upgradedBy != NULL); + lock->upgradedBy = NULL; + LWLockRelease(partitionLock); + } } else LWLockRelease(partitionLock); @@ -1322,6 +1416,7 @@ SetupLockInTable(LockMethod lockMethodTable, PGPROC *proc, lock->nGranted = 0; MemSet(lock->requested, 0, sizeof(int) * MAX_LOCKMODES); MemSet(lock->granted, 0, sizeof(int) * MAX_LOCKMODES); + lock->upgradedBy = NULL; LOCK_PRINT("LockAcquire: new", lock, lockmode); } else @@ -1768,6 +1863,12 @@ CleanUpLock(LOCK *lock, PROCLOCK *proclock, elog(PANIC, "proclock table corrupted"); } + /* + * Was this backend upgrading the lock? + */ + if (lock->upgradedBy == MyProc) + lock->upgradedBy = NULL; + if (lock->nRequested == 0) { /* diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c index 1ac25068d62..7b31370db70 100644 --- a/src/backend/storage/lmgr/proc.c +++ b/src/backend/storage/lmgr/proc.c @@ -95,7 +95,6 @@ static volatile sig_atomic_t got_deadlock_timeout; static void RemoveProcFromArray(int code, Datum arg); static void ProcKill(int code, Datum arg); static void AuxiliaryProcKill(int code, Datum arg); -static DeadLockState CheckDeadLock(void); /* @@ -1143,7 +1142,8 @@ AuxiliaryPidGetProc(int pid) * NOTES: The process queue is now a priority queue for locking. */ ProcWaitStatus -JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) +JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait, + bool isUpgrade) { LOCKMODE lockmode = locallock->tag.mode; LOCK *lock = locallock->lock; @@ -1230,8 +1230,12 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait) /* Must he wait for me? */ if (lockMethodTable->conflictTab[proc->waitLockMode] & myHeldLocks) { - /* Must I wait for him ? */ - if (lockMethodTable->conflictTab[lockmode] & proc->heldLocks) + /* + * Must I wait for him? I don't want a deadlock during lock + * upgrade - other processes should fail on it. + */ + if ((lockMethodTable->conflictTab[lockmode] & proc->heldLocks) && + !isUpgrade) { /* * Yes, so we have a deadlock. Easiest way to clean up @@ -1461,8 +1465,22 @@ ProcSleep(LOCALLOCK *locallock) (void) WaitLatch(MyLatch, WL_LATCH_SET | WL_EXIT_ON_PM_DEATH, 0, PG_WAIT_LOCK | locallock->tag.lock.locktag_type); ResetLatch(MyLatch); - /* check for deadlocks first, as that's probably log-worthy */ - if (got_deadlock_timeout) + /* + * Check for deadlocks first, as that's probably log-worthy. + * + * Do not wait for the timeout if the lock is being upgraded since + * the risk of deadlock is higher now. However, do not check for + * deadlock if the lock is being upgraded by this process - other + * processes should take care. + * + * TODO Possible optimization: if this is the only lock of the + * backend and if it did not have any weaker lock on the table so + * far, it should be safe to skip the deadlock check. However, to + * evaluate the situation, we need to take fast-path locks into + * account. + */ + if ((got_deadlock_timeout || lock->upgradedBy) && + lock->upgradedBy != MyProc) { deadlock_state = CheckDeadLock(); got_deadlock_timeout = false; @@ -1819,7 +1837,7 @@ ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock) * not, just return. If we have a real deadlock, remove ourselves from the * lock's wait queue. */ -static DeadLockState +DeadLockState CheckDeadLock(void) { int i; diff --git a/src/include/storage/lmgr.h b/src/include/storage/lmgr.h index 2a985ce5e15..eddbe1e8622 100644 --- a/src/include/storage/lmgr.h +++ b/src/include/storage/lmgr.h @@ -38,6 +38,7 @@ extern void RelationInitLockInfo(Relation relation); /* Lock a relation */ extern void LockRelationOid(Oid relid, LOCKMODE lockmode); +extern void LockRelationOidUpgrade(Oid relid, LOCKMODE lockmode); extern void LockRelationId(LockRelId *relid, LOCKMODE lockmode); extern bool ConditionalLockRelationOid(Oid relid, LOCKMODE lockmode); extern void UnlockRelationId(LockRelId *relid, LOCKMODE lockmode); diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h index ee3cb1dc203..080b69eddfd 100644 --- a/src/include/storage/lock.h +++ b/src/include/storage/lock.h @@ -131,6 +131,7 @@ typedef const LockMethodData *LockMethod; * nRequested -- total requested locks of all types. * granted -- count of each lock type currently granted on the lock. * nGranted -- total granted locks of all types. + * upgrader -- process that intends to upgrade the lock * * Note: these counts count 1 for each backend. Internally to a backend, * there may be multiple grabs on a particular lock, but this is not reflected @@ -150,6 +151,7 @@ typedef struct LOCK int nRequested; /* total of requested[] array */ int granted[MAX_LOCKMODES]; /* counts of granted locks */ int nGranted; /* total of granted[] array */ + PGPROC *upgradedBy; /* is lock being upgraded by this process? */ } LOCK; #define LOCK_LOCKMETHOD(lock) ((LOCKMETHODID) (lock).tag.locktag_lockmethodid) @@ -390,7 +392,8 @@ extern LockAcquireResult LockAcquireExtended(const LOCKTAG *locktag, bool dontWait, bool reportMemoryError, LOCALLOCK **locallockp, - bool logLockFailure); + bool logLockFailure, + bool isUpgrade); extern void AbortStrongLockAcquire(void); extern void MarkLockClear(LOCALLOCK *locallock); extern bool LockRelease(const LOCKTAG *locktag, diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 3e1d1fad5f9..c4bc3a24044 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -563,10 +563,12 @@ extern bool HaveNFreeProcs(int n, int *nfree); extern void ProcReleaseLocks(bool isCommit); extern ProcWaitStatus JoinWaitQueue(LOCALLOCK *locallock, - LockMethod lockMethodTable, bool dontWait); + LockMethod lockMethodTable, bool dontWait, + bool isUpgrade); extern ProcWaitStatus ProcSleep(LOCALLOCK *locallock); extern void ProcWakeup(PGPROC *proc, ProcWaitStatus waitStatus); extern void ProcLockWakeup(LockMethod lockMethodTable, LOCK *lock); +extern DeadLockState CheckDeadLock(void); extern void CheckDeadLockAlert(void); extern void LockErrorCleanup(void); extern void GetLockHoldersAndWaiters(LOCALLOCK *locallock, -- 2.47.3 --=-=-=-- ^ permalink raw reply [nested|flat] 956+ messages in thread
* [PATCH v6 3/4] psql: bump minimum supported version to v10 @ 2026-04-17 18:34 Nathan Bossart <nathan@postgresql.org> 0 siblings, 0 replies; 956+ messages in thread From: Nathan Bossart @ 2026-04-17 18:34 UTC (permalink / raw) --- doc/src/sgml/ref/psql-ref.sgml | 2 +- src/bin/psql/command.c | 23 +--- src/bin/psql/describe.c | 244 +-------------------------------- src/bin/psql/tab-complete.in.c | 46 +++---- 4 files changed, 27 insertions(+), 288 deletions(-) diff --git a/doc/src/sgml/ref/psql-ref.sgml b/doc/src/sgml/ref/psql-ref.sgml index 7c05afd4719..56c2692e618 100644 --- a/doc/src/sgml/ref/psql-ref.sgml +++ b/doc/src/sgml/ref/psql-ref.sgml @@ -5523,7 +5523,7 @@ PSQL_EDITOR_LINENUMBER_ARG='--line ' or an older major version. Backslash commands are particularly likely to fail if the server is of a newer version than <application>psql</application> itself. However, backslash commands of the <literal>\d</literal> family should - work with servers of versions back to 9.2, though not necessarily with + work with servers of versions back to 10, though not necessarily with servers newer than <application>psql</application> itself. The general functionality of running SQL commands and displaying query results should also work with servers of a newer major version, but this cannot diff --git a/src/bin/psql/command.c b/src/bin/psql/command.c index 01b8f11aadd..44ef11d980e 100644 --- a/src/bin/psql/command.c +++ b/src/bin/psql/command.c @@ -4471,10 +4471,10 @@ connection_warnings(bool in_startup) /* * Warn if server's major version is newer than ours, or if server - * predates our support cutoff (currently 9.2). + * predates our support cutoff (currently 10). */ if (pset.sversion / 100 > client_ver / 100 || - pset.sversion < 90200) + pset.sversion < 100000) printf(_("WARNING: %s major version %s, server major version %s.\n" " Some psql features might not work.\n"), pset.progname, @@ -6278,15 +6278,13 @@ get_create_object_cmd(EditableObjectType obj_type, Oid oid, * ensure the right view gets replaced. Also, check relation kind * to be sure it's a view. * - * Starting with PG 9.4, views may have WITH [LOCAL|CASCADED] + * Views may have WITH [LOCAL|CASCADED] * CHECK OPTION. These are not part of the view definition * returned by pg_get_viewdef() and so need to be retrieved - * separately. Materialized views (introduced in 9.3) may have + * separately. Materialized views may have * arbitrary storage parameter reloptions. */ printfPQExpBuffer(query, "/* %s */\n", _("Get view's definition and details")); - if (pset.sversion >= 90400) - { appendPQExpBuffer(query, "SELECT nspname, relname, relkind, " "pg_catalog.pg_get_viewdef(c.oid, true), " @@ -6297,19 +6295,6 @@ get_create_object_cmd(EditableObjectType obj_type, Oid oid, "LEFT JOIN pg_catalog.pg_namespace n " "ON c.relnamespace = n.oid WHERE c.oid = %u", oid); - } - else - { - appendPQExpBuffer(query, - "SELECT nspname, relname, relkind, " - "pg_catalog.pg_get_viewdef(c.oid, true), " - "c.reloptions AS reloptions, " - "NULL AS checkoption " - "FROM pg_catalog.pg_class c " - "LEFT JOIN pg_catalog.pg_namespace n " - "ON c.relnamespace = n.oid WHERE c.oid = %u", - oid); - } break; } diff --git a/src/bin/psql/describe.c b/src/bin/psql/describe.c index af3935b0078..ce99137c613 100644 --- a/src/bin/psql/describe.c +++ b/src/bin/psql/describe.c @@ -3,9 +3,9 @@ * * Support for the various \d ("describe") commands. Note that the current * expectation is that all functions in this file will succeed when working - * with servers of versions 9.2 and up. It's okay to omit irrelevant + * with servers of versions 10 and up. It's okay to omit irrelevant * information for an old server, but not to fail outright. (But failing - * against a pre-9.2 server is allowed.) + * against a pre-10 server is allowed.) * * Copyright (c) 2000-2026, PostgreSQL Global Development Group * @@ -154,16 +154,6 @@ describeAccessMethods(const char *pattern, bool verbose) printQueryOpt myopt = pset.popt; static const bool translate_columns[] = {false, true, false, false}; - if (pset.sversion < 90600) - { - char sverbuf[32]; - - pg_log_error("The server (version %s) does not support access methods.", - formatPGVersionNumber(pset.sversion, false, - sverbuf, sizeof(sverbuf))); - return true; - } - initPQExpBuffer(&buf); printfPQExpBuffer(&buf, "/* %s */\n", _("Get matching access methods")); @@ -312,9 +302,6 @@ describeFunctions(const char *functypes, const char *func_pattern, printQueryOpt myopt = pset.popt; static const bool translate_columns[] = {false, false, false, false, true, true, true, false, true, true, false, false, false, false}; - /* No "Parallel" column before 9.6 */ - static const bool translate_columns_pre_96[] = {false, false, false, false, true, true, false, true, true, false, false, false, false}; - if (strlen(functypes) != strspn(functypes, df_options)) { pg_log_error("\\df only takes [%s] as options", df_options); @@ -400,7 +387,6 @@ describeFunctions(const char *functypes, const char *func_pattern, gettext_noop("stable"), gettext_noop("volatile"), gettext_noop("Volatility")); - if (pset.sversion >= 90600) appendPQExpBuffer(&buf, ",\n CASE\n" " WHEN p.proparallel = " @@ -613,16 +599,8 @@ describeFunctions(const char *functypes, const char *func_pattern, myopt.title = _("List of functions"); myopt.translate_header = true; - if (pset.sversion >= 90600) - { myopt.translate_columns = translate_columns; myopt.n_translate_columns = lengthof(translate_columns); - } - else - { - myopt.translate_columns = translate_columns_pre_96; - myopt.n_translate_columns = lengthof(translate_columns_pre_96); - } printQuery(res, &myopt, pset.queryFout, false, pset.logfile); @@ -1108,38 +1086,6 @@ permissionsList(const char *pattern, bool showSystem) " ), E'\\n') AS \"%s\"", gettext_noop("Column privileges")); - if (pset.sversion >= 90500 && pset.sversion < 100000) - appendPQExpBuffer(&buf, - ",\n pg_catalog.array_to_string(ARRAY(\n" - " SELECT polname\n" - " || CASE WHEN polcmd != '*' THEN\n" - " E' (' || polcmd::pg_catalog.text || E'):'\n" - " ELSE E':'\n" - " END\n" - " || CASE WHEN polqual IS NOT NULL THEN\n" - " E'\\n (u): ' || pg_catalog.pg_get_expr(polqual, polrelid)\n" - " ELSE E''\n" - " END\n" - " || CASE WHEN polwithcheck IS NOT NULL THEN\n" - " E'\\n (c): ' || pg_catalog.pg_get_expr(polwithcheck, polrelid)\n" - " ELSE E''\n" - " END" - " || CASE WHEN polroles <> '{0}' THEN\n" - " E'\\n to: ' || pg_catalog.array_to_string(\n" - " ARRAY(\n" - " SELECT rolname\n" - " FROM pg_catalog.pg_roles\n" - " WHERE oid = ANY (polroles)\n" - " ORDER BY 1\n" - " ), E', ')\n" - " ELSE E''\n" - " END\n" - " FROM pg_catalog.pg_policy pol\n" - " WHERE polrelid = c.oid), E'\\n')\n" - " AS \"%s\"", - gettext_noop("Policies")); - - if (pset.sversion >= 100000) appendPQExpBuffer(&buf, ",\n pg_catalog.array_to_string(ARRAY(\n" " SELECT polname\n" @@ -1666,7 +1612,7 @@ describeOneTableDetails(const char *schemaname, : "''"), oid); } - else if (pset.sversion >= 100000) + else { appendPQExpBuffer(&buf, "SELECT c.relchecks, c.relkind, c.relhasindex, c.relhasrules, " @@ -1683,57 +1629,6 @@ describeOneTableDetails(const char *schemaname, : "''"), oid); } - else if (pset.sversion >= 90500) - { - appendPQExpBuffer(&buf, - "SELECT c.relchecks, c.relkind, c.relhasindex, c.relhasrules, " - "c.relhastriggers, c.relrowsecurity, c.relforcerowsecurity, " - "c.relhasoids, false as relispartition, %s, c.reltablespace, " - "CASE WHEN c.reloftype = 0 THEN '' ELSE c.reloftype::pg_catalog.regtype::pg_catalog.text END, " - "c.relpersistence, c.relreplident\n" - "FROM pg_catalog.pg_class c\n " - "LEFT JOIN pg_catalog.pg_class tc ON (c.reltoastrelid = tc.oid)\n" - "WHERE c.oid = '%s';", - (verbose ? - "pg_catalog.array_to_string(c.reloptions || " - "array(select 'toast.' || x from pg_catalog.unnest(tc.reloptions) x), ', ')\n" - : "''"), - oid); - } - else if (pset.sversion >= 90400) - { - appendPQExpBuffer(&buf, - "SELECT c.relchecks, c.relkind, c.relhasindex, c.relhasrules, " - "c.relhastriggers, false, false, c.relhasoids, " - "false as relispartition, %s, c.reltablespace, " - "CASE WHEN c.reloftype = 0 THEN '' ELSE c.reloftype::pg_catalog.regtype::pg_catalog.text END, " - "c.relpersistence, c.relreplident\n" - "FROM pg_catalog.pg_class c\n " - "LEFT JOIN pg_catalog.pg_class tc ON (c.reltoastrelid = tc.oid)\n" - "WHERE c.oid = '%s';", - (verbose ? - "pg_catalog.array_to_string(c.reloptions || " - "array(select 'toast.' || x from pg_catalog.unnest(tc.reloptions) x), ', ')\n" - : "''"), - oid); - } - else - { - appendPQExpBuffer(&buf, - "SELECT c.relchecks, c.relkind, c.relhasindex, c.relhasrules, " - "c.relhastriggers, false, false, c.relhasoids, " - "false as relispartition, %s, c.reltablespace, " - "CASE WHEN c.reloftype = 0 THEN '' ELSE c.reloftype::pg_catalog.regtype::pg_catalog.text END, " - "c.relpersistence\n" - "FROM pg_catalog.pg_class c\n " - "LEFT JOIN pg_catalog.pg_class tc ON (c.reltoastrelid = tc.oid)\n" - "WHERE c.oid = '%s';", - (verbose ? - "pg_catalog.array_to_string(c.reloptions || " - "array(select 'toast.' || x from pg_catalog.unnest(tc.reloptions) x), ', ')\n" - : "''"), - oid); - } res = PSQLexec(buf.data); if (!res) @@ -1761,8 +1656,7 @@ describeOneTableDetails(const char *schemaname, tableinfo.reloftype = (strcmp(PQgetvalue(res, 0, 11), "") != 0) ? pg_strdup(PQgetvalue(res, 0, 11)) : NULL; tableinfo.relpersistence = *(PQgetvalue(res, 0, 12)); - tableinfo.relreplident = (pset.sversion >= 90400) ? - *(PQgetvalue(res, 0, 13)) : 'd'; + tableinfo.relreplident = *(PQgetvalue(res, 0, 13)); if (pset.sversion >= 120000) tableinfo.relam = PQgetisnull(res, 0, 14) ? NULL : pg_strdup(PQgetvalue(res, 0, 14)); @@ -1781,8 +1675,6 @@ describeOneTableDetails(const char *schemaname, char *footers[3] = {NULL, NULL, NULL}; printfPQExpBuffer(&buf, "/* %s */\n", _("Get sequence information")); - if (pset.sversion >= 100000) - { appendPQExpBuffer(&buf, "SELECT pg_catalog.format_type(seqtypid, NULL) AS \"%s\",\n" " seqstart AS \"%s\",\n" @@ -1804,30 +1696,6 @@ describeOneTableDetails(const char *schemaname, "FROM pg_catalog.pg_sequence\n" "WHERE seqrelid = '%s';", oid); - } - else - { - appendPQExpBuffer(&buf, - "SELECT 'bigint' AS \"%s\",\n" - " start_value AS \"%s\",\n" - " min_value AS \"%s\",\n" - " max_value AS \"%s\",\n" - " increment_by AS \"%s\",\n" - " CASE WHEN is_cycled THEN '%s' ELSE '%s' END AS \"%s\",\n" - " cache_value AS \"%s\"\n", - gettext_noop("Type"), - gettext_noop("Start"), - gettext_noop("Minimum"), - gettext_noop("Maximum"), - gettext_noop("Increment"), - gettext_noop("yes"), - gettext_noop("no"), - gettext_noop("Cycles?"), - gettext_noop("Cache")); - appendPQExpBuffer(&buf, "FROM %s", fmtId(schemaname)); - /* must be separate because fmtId isn't reentrant */ - appendPQExpBuffer(&buf, ".%s;", fmtId(relationname)); - } res = PSQLexec(buf.data); if (!res) @@ -2045,10 +1913,7 @@ describeOneTableDetails(const char *schemaname, appendPQExpBufferStr(&buf, ",\n (SELECT c.collname FROM pg_catalog.pg_collation c, pg_catalog.pg_type t\n" " WHERE c.oid = a.attcollation AND t.oid = a.atttypid AND a.attcollation <> t.typcollation) AS attcollation"); attcoll_col = cols++; - if (pset.sversion >= 100000) appendPQExpBufferStr(&buf, ",\n a.attidentity"); - else - appendPQExpBufferStr(&buf, ",\n ''::pg_catalog.char AS attidentity"); attidentity_col = cols++; if (pset.sversion >= 120000) appendPQExpBufferStr(&buf, ",\n a.attgenerated"); @@ -2461,10 +2326,7 @@ describeOneTableDetails(const char *schemaname, CppAsString2(CONSTRAINT_EXCLUSION) ") AND " "condeferred) AS condeferred,\n"); - if (pset.sversion >= 90400) appendPQExpBufferStr(&buf, "i.indisreplident,\n"); - else - appendPQExpBufferStr(&buf, "false AS indisreplident,\n"); if (pset.sversion >= 150000) appendPQExpBufferStr(&buf, "i.indnullsnotdistinct,\n"); @@ -2569,10 +2431,7 @@ describeOneTableDetails(const char *schemaname, "pg_catalog.pg_get_indexdef(i.indexrelid, 0, true),\n " "pg_catalog.pg_get_constraintdef(con.oid, true), " "contype, condeferrable, condeferred"); - if (pset.sversion >= 90400) appendPQExpBufferStr(&buf, ", i.indisreplident"); - else - appendPQExpBufferStr(&buf, ", false AS indisreplident"); appendPQExpBufferStr(&buf, ", c2.reltablespace"); if (pset.sversion >= 180000) appendPQExpBufferStr(&buf, ", con.conperiod"); @@ -2823,17 +2682,11 @@ describeOneTableDetails(const char *schemaname, PQclear(result); /* print any row-level policies */ - if (pset.sversion >= 90500) - { printfPQExpBuffer(&buf, "/* %s */\n", _("Get row-level policies for this table")); appendPQExpBufferStr(&buf, "SELECT pol.polname,"); - if (pset.sversion >= 100000) appendPQExpBufferStr(&buf, " pol.polpermissive,\n"); - else - appendPQExpBufferStr(&buf, - " 't' as polpermissive,\n"); appendPQExpBuffer(&buf, " CASE WHEN pol.polroles = '{0}' THEN NULL ELSE pg_catalog.array_to_string(array(select rolname from pg_catalog.pg_roles where oid = any (pol.polroles) order by 1),',') END,\n" " pg_catalog.pg_get_expr(pol.polqual, pol.polrelid),\n" @@ -2904,7 +2757,6 @@ describeOneTableDetails(const char *schemaname, printTableAddFooter(&cont, buf.data); } PQclear(result); - } /* print any extended statistics */ if (pset.sversion >= 140000) @@ -3007,7 +2859,7 @@ describeOneTableDetails(const char *schemaname, } PQclear(result); } - else if (pset.sversion >= 100000) + else { printfPQExpBuffer(&buf, "/* %s */\n", _("Get extended statistics for this table")); @@ -3173,8 +3025,6 @@ describeOneTableDetails(const char *schemaname, } /* print any publications */ - if (pset.sversion >= 100000) - { printfPQExpBuffer(&buf, "/* %s */\n", _("Get publications that publish this table")); if (pset.sversion >= 150000) @@ -3284,7 +3134,6 @@ describeOneTableDetails(const char *schemaname, printTableAddFooter(&cont, buf.data); } PQclear(result); - } /* Print publications where the table is in the EXCEPT clause */ if (pset.sversion >= 190000) @@ -3706,7 +3555,7 @@ describeOneTableDetails(const char *schemaname, "ORDER BY pg_catalog.pg_get_expr(c.relpartbound, c.oid) = 'DEFAULT'," " c.oid::pg_catalog.regclass::pg_catalog.text;", oid); - else if (pset.sversion >= 100000) + else appendPQExpBuffer(&buf, "SELECT c.oid::pg_catalog.regclass, c.relkind," " false AS inhdetachpending," @@ -3716,14 +3565,6 @@ describeOneTableDetails(const char *schemaname, "ORDER BY pg_catalog.pg_get_expr(c.relpartbound, c.oid) = 'DEFAULT'," " c.oid::pg_catalog.regclass::pg_catalog.text;", oid); - else - appendPQExpBuffer(&buf, - "SELECT c.oid::pg_catalog.regclass, c.relkind," - " false AS inhdetachpending, NULL\n" - "FROM pg_catalog.pg_class c, pg_catalog.pg_inherits i\n" - "WHERE c.oid = i.inhrelid AND i.inhparent = '%s'\n" - "ORDER BY c.oid::pg_catalog.regclass::pg_catalog.text;", - oid); result = PSQLexec(buf.data); if (!result) @@ -3964,11 +3805,7 @@ describeRoles(const char *pattern, bool verbose, bool showSystem) ncols++; } appendPQExpBufferStr(&buf, "\n, r.rolreplication"); - - if (pset.sversion >= 90500) - { appendPQExpBufferStr(&buf, "\n, r.rolbypassrls"); - } appendPQExpBufferStr(&buf, "\nFROM pg_catalog.pg_roles r\n"); @@ -4023,7 +3860,6 @@ describeRoles(const char *pattern, bool verbose, bool showSystem) if (strcmp(PQgetvalue(res, i, (verbose ? 9 : 8)), "t") == 0) add_role_attribute(&buf, _("Replication")); - if (pset.sversion >= 90500) if (strcmp(PQgetvalue(res, i, (verbose ? 10 : 9)), "t") == 0) add_role_attribute(&buf, _("Bypass RLS")); @@ -4514,19 +4350,6 @@ listPartitionedTables(const char *reltypes, const char *pattern, bool verbose) const char *tabletitle; bool mixed_output = false; - /* - * Note: Declarative table partitioning is only supported as of Pg 10.0. - */ - if (pset.sversion < 100000) - { - char sverbuf[32]; - - pg_log_error("The server (version %s) does not support declarative table partitioning.", - formatPGVersionNumber(pset.sversion, false, - sverbuf, sizeof(sverbuf))); - return true; - } - /* If no relation kind was selected, show them all */ if (!showTables && !showIndexes) showTables = showIndexes = true; @@ -5034,16 +4857,6 @@ listEventTriggers(const char *pattern, bool verbose) static const bool translate_columns[] = {false, false, false, true, false, false, false}; - if (pset.sversion < 90300) - { - char sverbuf[32]; - - pg_log_error("The server (version %s) does not support event triggers.", - formatPGVersionNumber(pset.sversion, false, - sverbuf, sizeof(sverbuf))); - return true; - } - initPQExpBuffer(&buf); printfPQExpBuffer(&buf, "/* %s */\n", _("Get matching event triggers")); @@ -5113,16 +4926,6 @@ listExtendedStats(const char *pattern, bool verbose) PGresult *res; printQueryOpt myopt = pset.popt; - if (pset.sversion < 100000) - { - char sverbuf[32]; - - pg_log_error("The server (version %s) does not support extended statistics.", - formatPGVersionNumber(pset.sversion, false, - sverbuf, sizeof(sverbuf))); - return true; - } - initPQExpBuffer(&buf); printfPQExpBuffer(&buf, "/* %s */\n", _("Get matching extended statistics")); @@ -5352,7 +5155,6 @@ listCollations(const char *pattern, bool verbose, bool showSystem) gettext_noop("Schema"), gettext_noop("Name")); - if (pset.sversion >= 100000) appendPQExpBuffer(&buf, " CASE c.collprovider " "WHEN " CppAsString2(COLLPROVIDER_DEFAULT) " THEN 'default' " @@ -5361,10 +5163,6 @@ listCollations(const char *pattern, bool verbose, bool showSystem) "WHEN " CppAsString2(COLLPROVIDER_ICU) " THEN 'icu' " "END AS \"%s\",\n", gettext_noop("Provider")); - else - appendPQExpBuffer(&buf, - " 'libc' AS \"%s\",\n", - gettext_noop("Provider")); appendPQExpBuffer(&buf, " c.collcollate AS \"%s\",\n" @@ -6688,16 +6486,6 @@ listPublications(const char *pattern) printQueryOpt myopt = pset.popt; static const bool translate_columns[] = {false, false, false, false, false, false, false, false, false, false}; - if (pset.sversion < 100000) - { - char sverbuf[32]; - - pg_log_error("The server (version %s) does not support publications.", - formatPGVersionNumber(pset.sversion, false, - sverbuf, sizeof(sverbuf))); - return true; - } - initPQExpBuffer(&buf); printfPQExpBuffer(&buf, "/* %s */\n", _("Get matching publications")); @@ -6835,16 +6623,6 @@ describePublications(const char *pattern) PQExpBufferData title; printTableContent cont; - if (pset.sversion < 100000) - { - char sverbuf[32]; - - pg_log_error("The server (version %s) does not support publications.", - formatPGVersionNumber(pset.sversion, false, - sverbuf, sizeof(sverbuf))); - return true; - } - has_pubsequence = (pset.sversion >= 190000); has_pubtruncate = (pset.sversion >= 110000); has_pubgencols = (pset.sversion >= 180000); @@ -7095,16 +6873,6 @@ describeSubscriptions(const char *pattern, bool verbose) false, false, false, false, false, false, false, false, false, false, false, false, false, false, false, false, false}; - if (pset.sversion < 100000) - { - char sverbuf[32]; - - pg_log_error("The server (version %s) does not support subscriptions.", - formatPGVersionNumber(pset.sversion, false, - sverbuf, sizeof(sverbuf))); - return true; - } - initPQExpBuffer(&buf); printfPQExpBuffer(&buf, "/* %s */\n", _("Get matching subscriptions")); diff --git a/src/bin/psql/tab-complete.in.c b/src/bin/psql/tab-complete.in.c index 46b9add0604..f4d842966bf 100644 --- a/src/bin/psql/tab-complete.in.c +++ b/src/bin/psql/tab-complete.in.c @@ -1209,6 +1209,18 @@ Keywords_for_list_of_owner_roles, "PUBLIC" " FROM pg_catalog.pg_timezone_names() "\ " WHERE pg_catalog.quote_literal(pg_catalog.lower(name)) LIKE pg_catalog.lower('%s')" +#define Query_for_list_of_publications \ +"SELECT pubname "\ +" FROM pg_catalog.pg_publication "\ +" WHERE pubname LIKE '%s'" + +#define Query_for_list_of_subscriptions \ +"SELECT s.subname "\ +" FROM pg_catalog.pg_subscription s, pg_catalog.pg_database d"\ +" WHERE s.subname LIKE '%s' "\ +" AND d.datname = pg_catalog.current_database() "\ +" AND s.subdbid = d.oid" + /* Privilege options shared between GRANT and REVOKE */ #define Privilege_options_of_grant_and_revoke \ "SELECT", "INSERT", "UPDATE", "DELETE", "TRUNCATE", "REFERENCES", "TRIGGER", \ @@ -1243,32 +1255,6 @@ Copy_common_options, "DEFAULT", "FORCE_NOT_NULL", "FORCE_NULL", "FREEZE", \ #define Copy_to_options \ Copy_common_options, "FORCE_QUOTE", "FORCE_ARRAY" -/* - * These object types were introduced later than our support cutoff of - * server version 9.2. We use the VersionedQuery infrastructure so that - * we don't send certain-to-fail queries to older servers. - */ - -static const VersionedQuery Query_for_list_of_publications[] = { - {100000, - " SELECT pubname " - " FROM pg_catalog.pg_publication " - " WHERE pubname LIKE '%s'" - }, - {0, NULL} -}; - -static const VersionedQuery Query_for_list_of_subscriptions[] = { - {100000, - " SELECT s.subname " - " FROM pg_catalog.pg_subscription s, pg_catalog.pg_database d " - " WHERE s.subname LIKE '%s' " - " AND d.datname = pg_catalog.current_database() " - " AND s.subdbid = d.oid" - }, - {0, NULL} -}; - /* Known command-starting keywords. */ static const char *const sql_commands[] = { "ABORT", "ALTER", "ANALYZE", "BEGIN", "CALL", "CHECKPOINT", "CLOSE", "CLUSTER", @@ -1346,7 +1332,7 @@ static const pgsql_thing_t words_after_create[] = { {"POLICY", NULL, NULL, NULL}, {"PROCEDURE", NULL, NULL, Query_for_list_of_procedures}, {"PROPERTY GRAPH", NULL, NULL, &Query_for_list_of_propgraphs}, - {"PUBLICATION", NULL, Query_for_list_of_publications}, + {"PUBLICATION", Query_for_list_of_publications}, {"ROLE", Query_for_list_of_roles}, {"ROUTINE", NULL, NULL, &Query_for_list_of_routines, NULL, THING_NO_CREATE}, {"RULE", "SELECT rulename FROM pg_catalog.pg_rules WHERE rulename LIKE '%s'"}, @@ -1354,7 +1340,7 @@ static const pgsql_thing_t words_after_create[] = { {"SEQUENCE", NULL, NULL, &Query_for_list_of_sequences}, {"SERVER", Query_for_list_of_servers}, {"STATISTICS", NULL, NULL, &Query_for_list_of_statistics}, - {"SUBSCRIPTION", NULL, Query_for_list_of_subscriptions}, + {"SUBSCRIPTION", Query_for_list_of_subscriptions}, {"SYSTEM", NULL, NULL, NULL, NULL, THING_NO_CREATE | THING_NO_DROP}, {"TABLE", NULL, NULL, &Query_for_list_of_tables}, {"TABLESPACE", Query_for_list_of_tablespaces}, @@ -5701,9 +5687,9 @@ match_previous_words(int pattern_id, else if (TailMatchesCS("\\dP*")) COMPLETE_WITH_SCHEMA_QUERY(Query_for_list_of_partitioned_relations); else if (TailMatchesCS("\\dRp*")) - COMPLETE_WITH_VERSIONED_QUERY(Query_for_list_of_publications); + COMPLETE_WITH_QUERY(Query_for_list_of_publications); else if (TailMatchesCS("\\dRs*")) - COMPLETE_WITH_VERSIONED_QUERY(Query_for_list_of_subscriptions); + COMPLETE_WITH_QUERY(Query_for_list_of_subscriptions); else if (TailMatchesCS("\\ds*")) COMPLETE_WITH_SCHEMA_QUERY(Query_for_list_of_sequences); else if (TailMatchesCS("\\dt*")) -- 2.50.1 (Apple Git-155) --aHl3HfeiUuO3zeNC Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename=v6-0004-run-pgindent-and-pgperltidy.patch ^ permalink raw reply [nested|flat] 956+ messages in thread
end of thread, other threads:[~2026-04-17 18:34 UTC | newest] Thread overview: 956+ messages (download: mbox mbox.gz follow: Atom feed) -- links below jump to the message on this page -- 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 13:13 [PATCH] Lock upgrade without deadlocks. Antonin Houska <ah@cybertec.at> 2026-04-17 18:34 [PATCH v6 3/4] psql: bump minimum supported version to v10 Nathan Bossart <nathan@postgresql.org>
This inbox is served by agora; see mirroring instructions for how to clone and mirror all data and code used for this inbox